Papers with mean average
Vote’n’Rank: Revision of Benchmarking with Social Choice Theory (2023.eacl-main)
Copied to clipboard
Mark Rofin, Vladislav Mikhailov, Mikhail Florinsky, Andrey Kravchenko, Tatiana Shavrina, Elena Tutubalina, Daniel Karabekyan, Ekaterina Artemova
| Challenge: | ML benchmarks have been criticized for their construct validity, fragility of the design and task choices. |
| Approach: | They propose a framework for ranking systems in multi-task benchmarks under the principles of the social choice theory and propose 'vote'n'rank' procedures are more robust than the mean average while being able to handle missing performance scores and determine conditions under which the system becomes the winner. |
| Outcome: | The proposed framework can be utilised to draw new insights on benchmarking in several ML sub-fields and identify the best-performing systems in research and development case studies. |